Papers by Sathish Reddy Indurthi
Beyond Monolithic Rewards: Hybrid Multi-Aspect Reward Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to optimize for multimodal learning use a single reward mechanism, but they lack confidence calibration across domains. |
| Approach: | They propose a hybrid reward and multi-aspect reward modeling framework that integrates model-based and rule-based reward paradigms for accuracy and confidence calibration. |
| Outcome: | The proposed model improves accuracy and confidence calibration across multimodal tasks and introduces a generalized length-penalty reward to stabilize training and improve performance. |
Cut to the Chase: A Context Zoom-in Network for Reading Comprehension (D18-1)
Copied to clipboard
| Challenge: | Recent deep-learning based models suffer from reasoning over long documents and do not trivially generalize to cases where the answer is not present as a span. |
| Approach: | They propose a novel context zoom-in network (ConZNet) that can skip through irrelevant parts of a document and generate an answer using only the relevant regions of text. |
| Outcome: | The proposed architecture outperforms state-of-the-art results by 12.62% (ROUGE-L) relative improvement on the recently proposed and challenging RC dataset ‘NarrativeQA’. |
Look Harder: A Neural Machine Translation Model with Hard Attention (P19-1)
Copied to clipboard
| Challenge: | Soft-attention based Neural Machine Translation models attend all the words in the source sequence for each target token, which makes them ineffective for long sequence translation. |
| Approach: | They propose a hard-attention based NMT model which selects a subset of source tokens for each target token to effectively handle long sequence translation. |
| Outcome: | The proposed model performs better on long sequences and achieves significant improvement on English-German and English-French translation tasks compared to soft-attention based models. |
Language Model Augmented Monotonic Attention for Simultaneous Translation (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing adaptive policies for simultaneous neural machine translation use monotonic attention to perform read/write decisions based on the partial source and target sequences. |
| Approach: | They propose a framework to aid monotonic attention with an external language model to improve its decisions. |
| Outcome: | The proposed approach improves on English-German and English-French translation tasks by using a language model. |
Cross-lingual Evaluation of Multilingual Text Generation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for multilingual text generation are limited by language and data leakage. |
| Approach: | They propose an annotation-free cross-lingual evaluation protocol for multilingual text generation . they first generate English references from the translated non-English inputs into English . |
| Outcome: | The proposed protocol shows a high correlation to the reference-based ROUGE metric in four languages on news text summarization. |
WPO: Enhancing RLHF with Weighted Preference Optimization (2024.emnlp-main)
Copied to clipboard
Wenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu
| Challenge: | Off-policy preference optimization suffers from a distributional gap between the policy used for data collection and the target policy, leading to suboptimal optimization. |
| Approach: | They propose a method to simulate on-policy learning with off-police preference data. |
| Outcome: | The proposed method outperforms Direct Preference Optimization (DPO) by up to 5.6% on Alpaca Eval 2 and MT-bench. |
MemoReader: Large-Scale Reading Comprehension through Neural Memory Controller (D18-1)
Copied to clipboard
| Challenge: | Existing approaches to machine reading comprehension are limited in understanding, up to a few paragraphs, failing to comprehend lengthy documents. |
| Approach: | They propose a deep neural network architecture to handle a long-range dependency in RC tasks. |
| Outcome: | The proposed method outperforms existing methods especially for lengthy documents. |
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets (2024.findings-emnlp)
Copied to clipboard
Sathish Reddy Indurthi, Wenxuan Zhou, Shamil Chollampatt, Ravi Agrawal, Kaiqiang Song, Lingxiao Zhao, Chenguang Zhu
| Challenge: | Advancements in Large Language Models (LLMs) have significantly enhanced instruction-following capabilities, but most IFT datasets are predominantly in English, limiting model performance in other languages. |
| Approach: | They propose a method for collecting multilingual IFT datasets that preserves linguistic naturalness and ensures prompt diversity. |
| Outcome: | Experiments show that LLMs fine-tuned using this method show significant improvements in generative and discriminative tasks. |
ContextCheck: Sentence-Level Faithfulness Verification with Context-Aware Disambiguation (2026.findings-acl)
Copied to clipboard
Yueqin Yin, Yaxi Li, Xin Liu, Xun Wang, Kaiqiang Song, Simin Ma, Shujian Liu, Sathish Reddy Indurthi, Haoyun Deng, Pengcheng He, Mingyuan Zhou, Song Wang
| Challenge: | Large language models often hallucinate, producing content that is factually incorrect or not grounded in the sources. |
| Approach: | They propose a framework for sentence-level faithfulness verification with context-aware disambiguation. |
| Outcome: | The proposed framework improves Macro F1 by over 10 points compared to baselines on three context-dependent datasets. |